Skip to content

Install coherent deployment plans with dataset-bound SDS identities - #749

Merged
zzylol merged 26 commits into
mainfrom
refactor/backend-plan-split
Sep 28, 2026
Merged

zzylol merged 26 commits into
mainfrom
refactor/backend-plan-split

Conversation

@zzylol

@zzylol zzylol commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Deployments need one coherent installed plan that binds queries and producers to stored state with known meaning. Routing fingerprints alone cannot distinguish equal metric expressions over different logical datasets, or represent two deployed outputs with the same semantics safely.

Before this PR

Backend configuration mixes plan concerns and does not establish the SDS identity contract from #737. A query binding cannot independently express the semantic computation and the authorized deployed output.

After this PR

Install a coherent PrecomputePlan, QueryPlan, transmission configuration and immutable catalog snapshot. Planner exports the persisted output's canonical typed dependency closure, including logical dataset identity; deployment binding connects it to a stored output.

Planner persisted-output semantics + logical dataset
    → SummaryDefinition / definition_id
    → deployed stored_output_id
    → installed writer and query bindings
    → validated stored records

Two KLL(latency) outputs named hot/rebuild may share a definition but are independently bound. Another dataset changes the definition; relocating the same dataset preserves it. Same-version restart restores valid state. A new plan version remains cold until fresh input is produced and never adopts previous-version payloads.

Catalog schema 6 and planning snapshot schema 3 establish this identity contract here. Typed Count/Rate support formerly in #771 is included because those semantics must remain distinct. Shared physical execution and its precompute execution design document continue in #774.

Obsolete installed-plan aliases, untyped projection decoders, the unused materialization wire adapter, and cross-version adoption metadata are removed. Recovery accepts only the current sidecar schema and propagates malformed or unsupported metadata errors. Producer partition rosters are explicit on the wire.

Dependency changes

  • asap_sketch_codec contains the envelope encoding/decoding helpers used by backend ingest, DDSketch/KLL accumulators and wire-format tests. It supports removing the asap-precompute-rs Collector runtime dependency; it adds no sketch algorithms. The workspace member and data-plane dependency are intentional. With the Collector dependency removed, its Sketchlib patch block is unused and is removed too.
  • Planner advances from cd7e9e0 to bccc837 because this implementation uses LogicalDatasetIdentity and SummarySemanticFragment::from_stored_output_in_dataset. Both APIs are absent at the old revision. All four Planner dependencies use the same immutable revision.

Validation and scope

The SDS implementation passed locally: 117 type-library tests, 431 control-plane library tests, 1,165 data-plane library tests, 8 HTTP API tests, 13 serving integration tests, the production-process restart/warm-up test, strict all-target Clippy and formatting.

Coverage includes dataset changes, endpoint relocation, independent hot/rebuild bindings, tampered definitions, writer/read consistency, same-version restart without re-ingestion, and new-version warm-up. Earlier downstream validation also passed Level 1, exhaustive synthetic Level 2 selection, 332 storage tests and 13 serving integration tests.

Discovery, calibration and workload replay now preserve version-3 dataset identity. Process fixtures use typed deployment configuration instead of removed flat aggregation documents. At head c11a540d, the complete local workspace run passes all 1,794 tests, including 18 compatibility process tests, four sketch oracle tests, the control-plane → data-plane process test, and cold-successor activation. The 81 Python tests, workspace formatting and strict all-target Clippy also pass. External-service opt-in tests were not enabled; this is not manual deployment verification.

PR-specific evidence directories and generated archives are removed from the repository. Test summaries belong here; raw logs remain local or in CI.

Ad-hoc discovery, cross-version adoption and production-cost validation remain out of scope. Restricted native configuration helpers support explicit imported-state fixtures; production Planner compilation requires dataset identity. No manual deployment or human approval is claimed.

@zzylol zzylol changed the title refactor: split installed maintenance DAGs from query execution refactor: split backend plans and remove Collector dependency Sep 19, 2026
@zzylol zzylol changed the title refactor: split backend plans and remove Collector dependency refactor: split backend plans, unpin Planner, and remove Collector dependency Sep 21, 2026
@zzylol zzylol changed the title refactor: split backend plans, unpin Planner, and remove Collector dependency refactor: split backend plans and bind SDS state slots Sep 21, 2026
@zzylol
zzylol marked this pull request as ready for review September 21, 2026 16:13
@zzylol
zzylol changed the base branch from main to docs/physical-plan-design September 22, 2026 00:37
@zzylol zzylol changed the title refactor: split backend plans and bind SDS state slots refactor: split backend plans and bind stored summary outputs Sep 22, 2026
zzylol added a commit that referenced this pull request Sep 22, 2026
Integrate PR #749 reader/writer bindings and maintenance projections while preserving data-partition worker ownership and plan-derived configuration. Keep selected and maintenance DAG schemas distinct and adapt shared-sink execution to the projected graph.
@zzylol
zzylol force-pushed the docs/physical-plan-design branch from b0d77ce to 5f1eebf Compare September 28, 2026 13:47
@zzylol
zzylol force-pushed the refactor/backend-plan-split branch from 12896bd to 70a8d88 Compare September 28, 2026 16:14
@zzylol
zzylol changed the base branch from docs/physical-plan-design to main September 28, 2026 16:14
@zzylol zzylol changed the title refactor: split backend plans and bind stored summary outputs Install coherent deployment plans with dataset-bound SDS identities Sep 28, 2026
zzylol and others added 15 commits September 28, 2026 17:27
CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@zzylol
zzylol merged commit bb4a820 into main Sep 28, 2026
1 check passed
zzylol added a commit that referenced this pull request Sep 28, 2026
…774)

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* refactor: split installed maintenance DAGs from query execution

* refactor: remove backend Collector dependency and normalize legacy DAGs

* test: restore whole-backend process coverage without Collector

* fix: migrate Planner main and reject legacy runtime artifacts

* test: send full sketch envelope in whole-backend E2E

* refactor: bind backend summary state through versioned SDS slots

* test: send complete sketch envelopes in process fixtures

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* refactor: align plan bindings with current Planner and SDS contract

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* refactor: align plan bindings with summary store v1

* fix: validate selected DAG provenance

* fix: retain neutral sketch codec dependencies when syncing main

* fix: align derived DAG validation with current schema versions

* fix: keep maintenance document version distinct from complete DAG version

* refactor: adopt costed Planner selection without legacy API adapters

* test: verify exact process routing for uncertified Planner candidates

* docs: remove redundant Planner selection contract

* deps: pin Planner bounded HLL confidence model

* test: size transmitted KLL state from a certified accuracy contract

* test: retain KLL collector capability when using theoretical confidence

* refactor: implement typed summary semantics and physical plan lowering

* docs: separate Planner physical computation from backend deployment

* refactor: name the backend orchestration entry DeploymentPlanCompiler

* Restack PR #770 with implementation before standalone acceptance

* chore: consume Planner physical precompute candidate interfaces

* chore: consume Planner materialization frontier enumeration

* chore: consume exact temporal ranking physical candidates

* chore: use shared Planner candidate winner selection

* chore: consume sparse counter shared readout contracts

* test: declare collector ranking fixture capabilities explicitly

* Use Planner counter-window candidate execution

* Use shared keyed-counter omission contract

* Use Planner exact-counter population omission

* Construct fixture evidence for the standalone operator foundation

* build: pin Planner exact-state scratch merge implementation

* build: pin shared finalized-pane reconstruction fix

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* build: align shared Planner dependencies with remote PR 462

* docs: separate bound SDS range lookup from state validation

* feat(sds): separate bound output routing from persisted semantic identity

* docs: state SDS migration responsibilities without stale implementation claims

* refactor(sds): keep semantic variants compact without changing wire format

* test: align evidence fixtures with selected exact count state

* docs: align precompute SDS description with semantic catalog

* test: bind imported-state fixtures to their actual semantic definition

* test: distinguish imported CMS transport from total-count planning

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* docs: separate Planner physical computation from backend deployment

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* docs: separate bound SDS range lookup from state validation

* docs: state SDS migration responsibilities without stale implementation claims

* docs: preserve design index after rebasing onto main

* docs: clarify SDS source identity and version-scoped recovery

* docs: illustrate SDS identity and recovery decisions

* feat: bind Planner dataset semantics to deployment input identity

* fix: recognize shared native batch encoding at dependency boundary

* test: reject pre-dataset planning snapshot versions

* refactor: version dataset-bound catalog and update empty-plan fixtures

* docs: describe dataset-bound planning and installation inputs

* docs: clarify candidate selection and deployment ownership

* refactor: split backend deployment plans and bind stored outputs

* fix: restrict foundation SDS recovery to the installed generation

* fix: allocate fresh physical series when a plan version changes

* test: record foundation rebase and recovery regression evidence

* fix: complete shared state encoding adoption at the dependency boundary

* style: satisfy workspace formatting after the compiler rename

* feat(sds): separate bound output routing from persisted semantic identity

* feat(sds): require dataset-bound Planner definitions and versioned catalogs

* refactor: implement typed summary semantics and physical plan lowering

* fix: validate final SDS semantics across Planner adapters and recovery

* test: verify final SDS identity installation, serving and restart

* test: align shared execution fixtures with final #749 SDS foundation

* docs: record shared execution directly after the SDS foundation

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Remove obsolete SDS wire aliases and persisted metadata migration

* fix(data-plane): initialize dataset_identity in the bootstrap ingest contract

IngestContract gained an optional `dataset_identity`, but the bootstrap
envelope in data_plane's entry point was not updated, so the binary failed to
compile with E0063 and took the whole workspace build with it.

The bootstrap envelope describes the state before any plan is installed, where
every other field is a placeholder, so there is no dataset to bind to yet;
`None` is the accurate value. A published plan supplies the identity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject obsolete state-column syntax in compiler roundtrip coverage

* Remove unused materialization wire adapter and propagate recovery errors

* Keep projection fixtures on the canonical typed encoding

* Remove superseded deployment API field aliases

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* Document obsolete-format removals and cleanup validation

* Record passing restacked plan, storage and serving checks

* Remove generated evaluation archives and keep validation summaries

* Normalize validation summary formatting

* Remove PR process reports and defer execution design to shared-runtime PR

* Keep discovery and calibration on dataset-bound snapshot version 3

* Use typed deployment configuration in process E2E fixtures

* Align process assertions with dataset-bound SDS and cold successor activation

* fix: execute typed Planner fragments and preserve terminal resource errors

* fix: bind exact integer samples without losing input types

* test: verify exact integer protocol input binding

* refactor: name per-boundary physical fragments explicitly

* fix: consume finalized Planner query candidate outputs

* test: require Planner filters for bound protocol vectors

* test: use Planner schema lifting module

* fix: bind Planner filters and finalized query results

* docs: describe bound Planner filter execution

* fix: retain state binding and sharing beneath explicit query readouts

* fix: pin Planner query finalization for every candidate entry point

* docs: distinguish query values from stored accumulator boundaries

* test: assert exact Count state beneath its query readout

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 29, 2026
…eads (#763)

* docs: remove redundant Planner selection contract

* deps: pin Planner bounded HLL confidence model

* test: size transmitted KLL state from a certified accuracy contract

* test: retain KLL collector capability when using theoretical confidence

* refactor: implement typed summary semantics and physical plan lowering

* docs: separate Planner physical computation from backend deployment

* refactor: name the backend orchestration entry DeploymentPlanCompiler

* Restack PR #770 with implementation before standalone acceptance

* Restack PR #763 with implementation before standalone acceptance

* chore: consume Planner physical precompute candidate interfaces

* chore: consume Planner materialization frontier enumeration

* chore: consume exact temporal ranking physical candidates

* chore: use shared Planner candidate winner selection

* chore: consume sparse counter shared readout contracts

* test: declare collector ranking fixture capabilities explicitly

* Use Planner counter-window candidate execution

* Use shared keyed-counter omission contract

* Use Planner exact-counter population omission

* Construct fixture evidence for the standalone operator foundation

* build: pin Planner exact-state scratch merge implementation

* build: pin shared finalized-pane reconstruction fix

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* build: align shared Planner dependencies with remote PR 462

* docs: separate bound SDS range lookup from state validation

* feat(sds): separate bound output routing from persisted semantic identity

* docs: state SDS migration responsibilities without stale implementation claims

* refactor(sds): keep semantic variants compact without changing wire format

* test: align evidence fixtures with selected exact count state

* docs: align precompute SDS description with semantic catalog

* test: bind imported-state fixtures to their actual semantic definition

* test: bind imported-state fixtures to their actual semantic definition

* test: distinguish imported CMS transport from total-count planning

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* docs: separate Planner physical computation from backend deployment

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* docs: separate bound SDS range lookup from state validation

* docs: state SDS migration responsibilities without stale implementation claims

* docs: preserve design index after rebasing onto main

* docs: clarify SDS source identity and version-scoped recovery

* docs: illustrate SDS identity and recovery decisions

* feat: bind Planner dataset semantics to deployment input identity

* fix: reject cross-version stored-state adoption

* fix: recognize shared native batch encoding at dependency boundary

* test: reject pre-dataset planning snapshot versions

* refactor: version dataset-bound catalog and update empty-plan fixtures

* docs: describe dataset-bound planning and installation inputs

* docs: clarify candidate selection and deployment ownership

* refactor: split backend deployment plans and bind stored outputs

* fix: restrict foundation SDS recovery to the installed generation

* fix: allocate fresh physical series when a plan version changes

* fix: reserve the persisted storage handle during recovery

* test: record foundation rebase and recovery regression evidence

* fix: complete shared state encoding adoption at the dependency boundary

* style: satisfy workspace formatting after the compiler rename

* feat(sds): separate bound output routing from persisted semantic identity

* feat(sds): require dataset-bound Planner definitions and versioned catalogs

* refactor: implement typed summary semantics and physical plan lowering

* fix: validate final SDS semantics across Planner adapters and recovery

* test: verify final SDS identity installation, serving and restart

* test: align shared execution fixtures with final #749 SDS foundation

* docs: record shared execution directly after the SDS foundation

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Remove obsolete SDS wire aliases and persisted metadata migration

* fix(data-plane): initialize dataset_identity in the bootstrap ingest contract

IngestContract gained an optional `dataset_identity`, but the bootstrap
envelope in data_plane's entry point was not updated, so the binary failed to
compile with E0063 and took the whole workspace build with it.

The bootstrap envelope describes the state before any plan is installed, where
every other field is a placeholder, so there is no dataset to bind to yet;
`None` is the accurate value. A published plan supplies the identity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject obsolete state-column syntax in compiler roundtrip coverage

* Remove unused materialization wire adapter and propagate recovery errors

* Keep projection fixtures on the canonical typed encoding

* Remove superseded deployment API field aliases

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* Document obsolete-format removals and cleanup validation

* Record passing restacked plan, storage and serving checks

* Remove generated evaluation archives and keep validation summaries

* Normalize validation summary formatting

* Remove PR process reports and defer execution design to shared-runtime PR

* Keep discovery and calibration on dataset-bound snapshot version 3

* Use typed deployment configuration in process E2E fixtures

* Remove unused streaming-config fixture after physical-only startup

* Align process assertions with dataset-bound SDS and cold successor activation

* fix: execute typed Planner fragments and preserve terminal resource errors

* fix: bind exact integer samples without losing input types

* test: verify exact integer protocol input binding

* refactor: name per-boundary physical fragments explicitly

* fix: consume finalized Planner query candidate outputs

* test: require Planner filters for bound protocol vectors

* test: use Planner schema lifting module

* fix: bind Planner filters and finalized query results

* docs: describe bound Planner filter execution

* fix: retain state binding and sharing beneath explicit query readouts

* fix: pin Planner query finalization for every candidate entry point

* docs: distinguish query values from stored accumulator boundaries

* test: assert exact Count state beneath its query readout

* fix: preserve maintenance failures and stop failed worker admission

* docs: specify bounded revisions and consistent query snapshots

* feat: execute bounded precompute revisions with consistent SDS reads

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 29, 2026
…765)

* refactor(sds): keep semantic variants compact without changing wire format

* test: align evidence fixtures with selected exact count state

* docs: align precompute SDS description with semantic catalog

* test: bind imported-state fixtures to their actual semantic definition

* test: bind imported-state fixtures to their actual semantic definition

* test: distinguish imported CMS transport from total-count planning

* docs: clarify Planner physical plan and SDS architecture

* docs: specify executable subplan materialization boundaries

* docs: scope migration to backend precompute and query plans

* docs: clarify window terminology migration

* Revert "docs: clarify window terminology migration"

This reverts commit daa5281.

* docs: focus physical plan and SDS designs

* docs: add physical compiler input example

* docs: add physical compiler output example

* docs: include query expression in compiler example

* docs: reorganize physical plan integration design

* docs: reorganize SDS and migration designs

* docs: define maintenance inputs before plan split example

* docs: use plan version consistently in backend design

* docs: remove standalone catalog materialization abstraction

* docs: add concise planner backend glossary

* docs: clarify selected deployment guarantee terminology

* docs: explain missing planner maintenance guarantee

* docs: motivate selected producer maintenance decision

* docs: label catalog reads and SDS metadata ownership

* docs: align SDS ownership and lifecycle terminology

* docs: use current planner and plan-version names consistently

* docs: align integration diagram with SDS ownership

* docs: clarify instance identity and shared producer wording

* docs: distinguish summary definitions from runtime stores

* docs: model one runtime summary store for DAG bindings

* docs: scope SDS lifecycle to read eligibility

* docs: tie stored summary examples directly to DAG outputs

* docs: name summary tables, stored records, and output references by role

* docs: illustrate summary definitions, stored records, and output references

* docs: limit v1 summary storage to definitions and stored summaries

* docs: separate Planner physical computation from backend deployment

* docs: bind Planner physical DAGs without backend re-lowering

* docs: describe summary inputs with groups and pane duration

* docs: define summary semantic completeness beyond input scope

* docs: define SDS identity through canonical Planner computation

* docs: decouple SDS semantic identity from executable Planner IR

* docs: define Planner-owned SDS discovery for future ad hoc queries

* docs: streamline SDS design around definitions and stored results

* docs: track bound-query SDS migration across implementation PRs

* docs: identify active shared-library PR in bound-query migration

* docs: separate bound SDS range lookup from state validation

* docs: state SDS migration responsibilities without stale implementation claims

* docs: preserve design index after rebasing onto main

* docs: clarify SDS source identity and version-scoped recovery

* docs: illustrate SDS identity and recovery decisions

* feat: bind Planner dataset semantics to deployment input identity

* fix: reject cross-version stored-state adoption

* fix: recognize shared native batch encoding at dependency boundary

* test: reject pre-dataset planning snapshot versions

* refactor: version dataset-bound catalog and update empty-plan fixtures

* test: enforce whole-query consistency and new-version warm-up

* docs: describe dataset-bound planning and installation inputs

* docs: clarify candidate selection and deployment ownership

* docs: align installation guide with dataset and recovery contracts

* refactor: split backend deployment plans and bind stored outputs

* fix: restrict foundation SDS recovery to the installed generation

* fix: allocate fresh physical series when a plan version changes

* fix: reserve the persisted storage handle during recovery

* test: record foundation rebase and recovery regression evidence

* fix: complete shared state encoding adoption at the dependency boundary

* style: satisfy workspace formatting after the compiler rename

* feat(sds): separate bound output routing from persisted semantic identity

* feat(sds): require dataset-bound Planner definitions and versioned catalogs

* refactor: implement typed summary semantics and physical plan lowering

* fix: validate final SDS semantics across Planner adapters and recovery

* test: verify final SDS identity installation, serving and restart

* test: align shared execution fixtures with final #749 SDS foundation

* docs: record shared execution directly after the SDS foundation

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Remove obsolete SDS wire aliases and persisted metadata migration

* fix(data-plane): initialize dataset_identity in the bootstrap ingest contract

IngestContract gained an optional `dataset_identity`, but the bootstrap
envelope in data_plane's entry point was not updated, so the binary failed to
compile with E0063 and took the whole workspace build with it.

The bootstrap envelope describes the state before any plan is installed, where
every other field is a placeholder, so there is no dataset to bind to yet;
`None` is the accurate value. A published plan supplies the identity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject obsolete state-column syntax in compiler roundtrip coverage

* Remove unused materialization wire adapter and propagate recovery errors

* Keep projection fixtures on the canonical typed encoding

* Remove superseded deployment API field aliases

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* fix(control-plane): supply dataset_identity in the API test fixtures

CompileAndPublishPhysicalPlanRequest gained a required `dataset_identity`
field in this branch, but two api_tests fixtures build their request body as
JSON by hand and were never updated, so both failed deserialization with
`missing field dataset_identity` before reaching the handler.

Take the value from the planning snapshot's own `environment.dataset_identity`
rather than inventing one, so the fixture keeps describing the same dataset the
rest of the snapshot describes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit e627dbb)

* Document obsolete-format removals and cleanup validation

* Record passing restacked plan, storage and serving checks

* Remove generated evaluation archives and keep validation summaries

* Normalize validation summary formatting

* Remove PR process reports and defer execution design to shared-runtime PR

* Keep discovery and calibration on dataset-bound snapshot version 3

* Use typed deployment configuration in process E2E fixtures

* Remove unused streaming-config fixture after physical-only startup

* Align process assertions with dataset-bound SDS and cold successor activation

* fix: execute typed Planner fragments and preserve terminal resource errors

* fix: bind exact integer samples without losing input types

* test: retain resource failures across summary DAG execution

* fix: treat cooperative query yields as pending execution

* test: verify exact integer protocol input binding

* refactor: name per-boundary physical fragments explicitly

* fix: consume finalized Planner query candidate outputs

* test: expose terminal error masking during revision changes

* test: require Planner filters for bound protocol vectors

* test: use Planner schema lifting module

* fix: bind Planner filters and finalized query results

* fix: preserve terminal execution errors before revision fencing

* docs: describe bound Planner filter execution

* refactor: return typed execution error directly

* fix: retain state binding and sharing beneath explicit query readouts

* fix: pin Planner query finalization for every candidate entry point

* docs: distinguish query values from stored accumulator boundaries

* test: assert exact Count state beneath its query readout

* fix: preserve maintenance failures and stop failed worker admission

* docs: specify bounded revisions and consistent query snapshots

* feat: execute bounded precompute revisions with consistent SDS reads

* fix: enforce query-wide resource and snapshot contracts

* test: reject deployment candidates without snapshot agreement

* docs: clarify local versus external query input candidates

* test: retain Float64 formatting in external-only SQL oracle

* docs: clarify precomputation terminology and remove stale review notes

* docs: structure query DAG design around ownership and execution contracts

* docs: align query execution layers with Planner physical candidates

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant